Ollama 0.19 switches the Apple Silicon backend to MLX, achieving 1,810 tokens/s prefill and 112 tokens/s decode. NVFP4 quantization support and cache improvements landed at the same time.
An update moved inference from FP16 to BF16, and M1-M3 Macs emulate BF16 at half speed. Benchmarks of the regression and the config that lands at 2:30.
Hypura breaks away from llama.cpp’s mmap design and streams even dense models with a three-tier NVMe placement, while TurboQuant eliminates quantization-constant overhead via a polar-coordinate transform. Includes a design comparison with Flash‑MoE and a review of scenarios where KV‑cache compression actually helps.
Local video generation test on M1 Max 64GB MacBook Pro. FP8 models don't work on Metal — switching to GGUF got Wan 2.2 running at 82 minutes for a 2-second clip. LTX-2 produced NaN or unusable KSampler output under MPS. Specs, failed configs, and the working setup.
Upscaling images loaded via the Load Image node was producing garbled output. Fixed it by addressing the non-contiguous tensor issue — a one-line patch to comfy/utils.py. Added a 2026-04-29 follow-up after a ComfyUI update wiped the patch and the bug came back, with the upstream PyTorch issue and a recurrence-detection snippet.
A derivative checkpoint of Z-Image Turbo released on ModelScope. It is tuned for skin texture and film-photography-like aesthetics, and can run on an M1 Max with 64GB.
A look at ACE-Step, the 'Stable Diffusion of music,' covering its architecture, features, installation, and expected performance on Apple Silicon before trying it on an M1 Max.
An overview of Z-Image-Distilled, the distilled fast-inference variant of Z-Image, including how it compares with FLUX.1 Schnell, how it runs on an M1 Max 64GB machine, and LoRA compatibility.
Klein 9B is undistilled, wants ~29GB, and RTX 4090 speeds don't translate to MPS. Where the CUDA gap comes from, and the 4B model that hits 30-40s on M1 Max.
A plan to build an internal help desk RAG system using a Mac mini M4 Pro and Dify. Highlights what's new in Dify circa 2025 and tips for running local LLMs.